Papers with text processing
An Empirical Study of Tokenization Strategies for Various Korean NLP Tasks (2020.aacl-main)
Copied to clipboard
| Challenge: | Traditionally, tokenization is the very first step in most text processing works. |
| Approach: | They propose to use morphological segmentation followed by BPE for Korean NLP tasks . they empirically examine what is the best tokenization strategy for Korean to/from English . |
| Outcome: | The proposed approach is best for Korean to/from English machine translation and natural language understanding tasks. |
Octopus: On-device language model for function calling of software APIs (2025.naacl-industry)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are pivotal for advanced text processing and generation. |
| Approach: | They propose a framework to train on-device Large Language Models optimized for invoking software APIs. |
| Outcome: | The proposed model outperforms GPT-4 in API calling tasks while maintaining inference speed. |
On-Device Neural Language Model Based Word Prediction (C18-2)
Copied to clipboard
| Challenge: | Currently, on-device keyboards have limited memory and response time for word prediction . a proposed on-device neural language model based word prediction method is available for mobile devices . |
| Approach: | They propose an on-device neural language model based word prediction method that optimizes run-time memory and provides a real-time prediction environment. |
| Outcome: | The proposed model outperforms existing methods for word prediction in keystroke savings and word prediction rate and has been commercialized. |
VISPool: Enhancing Transformer Encoders with Vector Visibility Graph Neural Networks (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing graph-based graph construction methods rely on static graphs and are not scalable with increasing document and word counts. |
| Approach: | They propose a dynamic graph construction method based on vector visibility graphs (VVGs) they propose scalable model architecture that integrates VVG convolutional networks into transformer pipelines. |
| Outcome: | The proposed model outperforms baseline models on the GLUE benchmark datasets. |
Finding the Law: Enhancing Statutory Article Retrieval via Graph Neural Networks (2023.eacl-main)
Copied to clipboard
| Challenge: | Statutory article retrieval (SAR) is a promising application of legal text processing. |
| Approach: | They propose a graph-augmented dense statute retriever model that incorporates the structure of legislation via a neural network to improve density retrieval performance. |
| Outcome: | The proposed model outperforms baselines on a real-world expert-annotated dataset. |
Learning Context-Sensitive Convolutional Filters for Text Processing (D18-1)
Copied to clipboard
| Challenge: | Convolutional neural networks (CNNs) are a popular building block for natural language processing . despite their success, most existing CNN models share the same learned set of filters for all input sentences. |
| Approach: | They propose to use a meta network to learn context-sensitive convolutional filters for text processing by using a bidirectional filter generation mechanism. |
| Outcome: | The proposed framework outperforms standard and attention-based CNN models on four different tasks. |
Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects (2026.acl-long)
Copied to clipboard
| Challenge: | Research on cross-dialectal transfer from a standard to a non-standard dialect variety has typically focused on text data. |
| Approach: | They compare standard-to-dialect transfer in three settings: text models, speech models, and cascaded systems where speech first gets automatically transcribed and then further processed by a text model. |
| Outcome: | The proposed model performs best on German dialect data while the text-only model perform best on the standard data. |
A unified approach to sentence segmentation of punctuated text in many languages (2021.acl-long)
Copied to clipboard
| Challenge: | Existing tools for segmenting punctuated text in many languages are limited in their language coverage and evaluation is ad hoc. |
| Approach: | They propose a new context-based modeling approach that can be trained on noisily-annotated data. |
| Outcome: | The proposed model exceeds baselines set by existing methods on English corpora and performs well on average on new multilingual evaluation set. |
Neural Topic Model with Reinforcement Learning (D19-1)
Copied to clipboard
| Challenge: | Experimental results show superior performance on perplexity and topic coherence measures compared to state-of-the-art topic models. |
| Approach: | They propose to incorporate topic coherence measures as reward signals to guide the learning of a VAE-based topic model. |
| Outcome: | The proposed model is able to separating background words dynamically from topic words eliminating the pre-processing step of filtering infrequent and/or top frequent words, typically required for learning traditional topic models. |
Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent years have witnessed significant advancements in large language models (LLMs) but still struggle with integrating vision and audio. |
| Approach: | They propose a self-knowledge distillation method to improve vision-audio capabilities of OLLMs by learning from the vision-text components. |
| Outcome: | The proposed method improves vision-audio capabilities of OLLMs by learning from vision-text components, which improves interaction between audio and images and results in improved performance on multimodal tasks. |
HIT - A Hierarchically Fused Deep Attention Network for Robust Code-mixed Language Representation (2021.findings-acl)
Copied to clipboard
| Challenge: | linguistics and morphology of resource-short code-mixed texts remain a key challenge in text processing. |
| Approach: | They propose a hierarchical transformer-based framework that captures the semantic relationship among words and hierarchically learns sentencelevel semantics using a fused attention mechanism. |
| Outcome: | The proposed framework improves on one European and five Indic languages on four NLP tasks on eleven datasets. |
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are generalist agents capable of operating within complex environments. |
| Approach: | They propose a class of tools that can serve as a middleware layer shielding LLMs from environmental complexity. |
| Outcome: | The proposed tool can shield the LLM from environmental complexity in two representative complex environments. |
Spherical Latent Spaces for Stable Variational Autoencoders (D18-1)
Copied to clipboard
| Challenge: | Variational autoencoders use a multivariate Gaussian latent variable to capture latent structure in data. |
| Approach: | They propose a variational autoencoder which uses a latent distribution instead of Gaussian . they find that the variational posterior averts the KL collapse by a fixed hyperparameter . |
| Outcome: | The von Mises-Fisher distribution averts the KL collapse and gives better likelihoods than Gaussian models across a range of modeling conditions. |
Multilingual Culture-Independent Word Analogy Datasets (2020.lrec-1)
Copied to clipboard
| Challenge: | In text processing, deep neural networks use word embeddings as an input. |
| Approach: | They propose to use benchmark datasets to compare the quality of word embeddings in text processing . they use a word analogy task in Croatian, English, Estonian, Finnish, Latvian, Lithuanian, Russian, Slovenian, and Swedish . |
| Outcome: | The proposed datasets are culturally independent and cross-lingual for the languages used. |
Sequential Learning of Convolutional Features for Effective Text Classification (D19-1)
Copied to clipboard
| Challenge: | Existing models for text classification have largely ignored convolution filters and max pooling . text classification is one of the major applications of natural language processing . |
| Approach: | They propose a convolutional attentive recurrent network model which uses convolution filters and max pooling to improve text classification. |
| Outcome: | The proposed model outperforms existing convolutional models on text classification tasks. |
Consonant is all you need: a compact representation of English text for efficient NLP (2023.findings-emnlp)
Copied to clipboard
| Challenge: | In natural language processing, the representation of text plays a crucial role in various tasks such as language modeling, sentiment analysis, and machine translation. |
| Approach: | They propose a method to represent English text with only consonants that is more discriminative than vowels and a technique to retrieve vowel information from it. |
| Outcome: | The proposed representation significantly reduces the overall memory and compute footprint required for storing and processing textual data. |
C²RBench: A Chinese Complex Reasoning Benchmark for Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks often fail to capture complex multi-step reasoning demands inherent in real-world scenarios. |
| Approach: | They propose a benchmark to evaluate multi-step, multimodal advanced reasoning of large language models. |
| Outcome: | The proposed benchmark exceeds existing benchmarks in cognitive complexity and accuracy by over 90% . it features 1,115 carefully curated Chinese tasks organized into eight domain-specific subsets . evaluations of 20 LLMs and 24 multimodal large language models reveal critical performance gaps . |
Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination (2026.acl-long)
Copied to clipboard
Lirong Gao, Zeqing Wang, Yuyan Cai, Jiayi Deng, Yanmei Gu, Yiming Zhang, Jia Zhou, Yanfei Zhang, Junbo Zhao
| Challenge: | Existing benchmarks assess basic knowledge breadth or lexical understanding, failing to capture higher-order skills that are central to historical research. |
| Approach: | They propose a benchmark anchored in the Chinese Imperial Examination system that assesses historical knowledge and lexical understanding. |
| Outcome: | The new benchmark aims to assess the ability of LLMs to process historical materials and documents. |